Conversation
Captures a post-0.1 direction (explicitly NOT a launch item): a third-party/ cloud compute backend and the storage substrate it forces. The through-line is that Slurm's shared filesystem was quietly doing three jobs — kernel state, resident<->job channel, and run output — and cloud has no cheap fleet-wide shared FS, so each gets a home that doesn't depend on a shared disk: - git-in / artifact-out instead of a region-pinned shared filesystem, so the substrate is region-agnostic and jobs can chase cheap capacity; a shared/ parallel FS, if any, is scoped inside one multi-node job in one AZ. - the resident decoupled from where experiments run (per-run backend choice; cloud VM or serverless cron; "job zero" bootstrap; drives Slurm + cloud at once). - durable state as a storage-backend choice (disk -> volume -> bucket), never a database — state stays blob-shaped and single-writer. - SkyPilot-first meta-provider so one backend covers many clouds. - a run-artifact store: results, reports, traces, and full measurement trajectories as artifacts keyed by run_id; a small metric map + provenance in the record. Traces private-by-default and bench-shaped; measurement standardized and optional; no metrics DB, W&B stays an optional live view. Roadmap pointer under Beyond 1.0; cross-references external.md and scaling.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Round 1 — reviewed head fc5e6c36 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: nothing blocking — 6 advisory notes.
6 findings attached to the lines below.
Merged verdict: the documented compute protocol is inaccurate; OUTERLOOP_ROOT is the state root; and job-local traces are not durable across preemption without checkpoint artifact uploads. Three prose rewrites are included verbatim as suggestions. Rejected: no substantive findings were rejected; repeated interface and trace-durability reports were deduplicated and retained with all agreeing lens attributions.
|
|
||
| ## The seam is already the right shape | ||
|
|
||
| A `compute` backend is about five operations: **submit** a job (resource spec + |
There was a problem hiding this comment.
The compute interface is described with operations it does not expose. [coverage+deployment+general+prose] Compute defines submit, status, pending_reason, job_partition, active_job_names, queue_snapshot, lane_load, job_id_for_name, and cancel—not fetch, poll, capacity, or placement. Rewrite it as: "A compute backend submits, reports job state and queue information, and cancels jobs; placement is chosen by the dispatch settings."
(high confidence)
| Durability is a **storage-backend choice**, in increasing weight: | ||
|
|
||
| - **Persistent volume (the default).** A cheap VM with an attached disk; the | ||
| state root (`OUTERLOOP_HOME`) is a mounted POSIX directory, exactly as on |
There was a problem hiding this comment.
The note names the checkout path as the state root. [deployment] The chain documents OUTERLOOP_ROOT as the required state root and OUTERLOOP_HOME as the checkout; rewrite "state root (OUTERLOOP_HOME)" as "state root (OUTERLOOP_ROOT)" to avoid mounting persistent state at the wrong path.
(high confidence)
| archival-science claim), age out aborted/negative ones after a window or | ||
| downsample. Same policy class as the run-dir reaper, moved to the bucket. | ||
|
|
||
| Because a trace streams to a local file and flushes at checkpoints + session end, |
There was a problem hiding this comment.
Preempted jobs lose locally flushed traces. [coverage+credentials+deployment+general+lifecycle] Traces flush only to job-local disk while the object-store artifact is sent at job end; preemption can destroy every checkpointed local trace on the ephemeral job. Specify checkpoint uploads to durable artifact storage or remove the crash-safety claim.
(high confidence)
| the `compute` interface and the storage interface — plus the metric schema | ||
| sketched in the metric-taxonomy work.** | ||
|
|
||
| The whole design turns on one observation: Slurm gave us a shared filesystem for |
There was a problem hiding this comment.
Suggestion. The opening explanation is padded and metaphorical. [prose] Rewrite it as: "Slurm shared filesystem stores kernel state, carries messages between the resident and jobs, and stores run output; cloud deployments need a separate service for each use."
(high confidence)
| - **Ephemeral per-job first** (spin up → run → tear down); a warm pool later if | ||
| cold-start latency bites — the same resident-vs-per-cadence choice we already | ||
| made on Slurm. | ||
| - **Teardown is load-bearing** for two reasons at once: you pay per second, and a |
There was a problem hiding this comment.
Suggestion. The teardown requirement uses a metaphor instead of stating the risk. [prose] Rewrite it as: "Teardown is required because cloud instances cost money while they run and may retain credentials after the job ends."
(high confidence)
|
|
||
| ## The run-artifact store | ||
|
|
||
| The final scalar was always an impoverishment. A run produces a **bundle of |
There was a problem hiding this comment.
Suggestion. The artifact-store introduction makes an unsupported grand claim. [prose] Rewrite it as: "A scalar result omits useful outputs, so each run should store results, reports, traces, and measurement trajectories as artifacts keyed by run_id."
(high confidence)
There was a problem hiding this comment.
Round 1 — reviewed head fc5e6c36 — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: nothing blocking — 1 advisory note.
1 finding attached to the lines below.
One advisory documentation finding: the described compute seam does not match the current interface.
|
|
||
| ## The seam is already the right shape | ||
|
|
||
| A `compute` backend is about five operations: **submit** a job (resource spec + |
There was a problem hiding this comment.
The compute interface does not provide the five operations listed here. exposes submit, status, cancellation, and queue helpers, while output is read directly from run directories and it has no fetch or capacity/placement method, so saying both current backends already implement this interface misstates the existing seam.
(high confidence)
What
A design sketch capturing tonight's brainstorm with Mengye on how Outerloop would run on third-party/cloud compute, and the storage substrate that forces. New file
docs/design/compute-cloud.md, a roadmap pointer under Beyond 1.0, and a CHANGELOG entry.Explicitly post-0.1 — not a launch item. v0.1 ships the two battle-tested backends (Slurm + LocalCompute). This note exists so the picture is ready when the cloud version comes up.
The through-line
Slurm's shared filesystem was quietly doing three jobs — kernel state, the resident↔job channel, and run output. Cloud has no cheap fleet-wide shared FS, so each gets a home that doesn't need a shared disk:
save_sample); measurement standardized + optional; no metrics DB; W&B stays an optional live view.Notes
external.md(which names the compute + storage interfaces at a high level) andscaling.md; lands the metric-taxonomy vocabulary as the measurement-record schema.🤖 Generated with Claude Code